Semi-parametric differential expression analysis via partial mixture estimation.
نویسندگان
چکیده
We develop an approach for microarray differential expression analysis, i.e. identifying genes whose expression levels differ between two or more groups. Current approaches to inference rely either on full parametric assumptions or on permutation-based techniques for sampling under the null distribution. In some situations, however, a full parametric model cannot be justified, or the sample size per group is too small for permutation methods to be valid. We propose a semi-parametric framework based on partial mixture estimation which only requires a parametric assumption for the null (equally expressed) distribution and can handle small sample sizes where permutation methods break down. We develop two novel improvements of Scott's minimum integrated square error criterion for partial mixture estimation [Scott, 2004a,b]. As a side benefit, we obtain interpretable and closed-form estimates for the proportion of EE genes. Pseudo-Bayesian and frequentist procedures for controlling the false discovery rate are given. Results from simulations and real datasets indicate that our approach can provide substantial advantages for small sample sizes over the SAM method of Tusher et al. [2001], the empirical Bayes procedure of Efron and Tibshirani [2002], the mixture of normals of Pan et al. [2003] and a t-test with p-value adjustment [Dudoit et al., 2003] to control the FDR [Benjamini and Hochberg, 1995].
منابع مشابه
Statistical Applications in Genetics and Molecular Biology
We develop an approach for microarray differential expression analysis, i.e. identifying genes whose expression levels differ between two or more groups. Current approaches to inference rely either on full parametric assumptions or on permutation-based techniques for sampling under the null distribution. In some situations, however, a full parametric model cannot be justified, or the sample siz...
متن کاملAn EM-based semi-parametric mixture model approach to the regression analysis of competing-risks data.
We consider a mixture model approach to the regression analysis of competing-risks data. Attention is focused on inference concerning the effects of factors on both the probability of occurrence and the hazard rate conditional on each of the failure types. These two quantities are specified in the mixture model using the logistic model and the proportional hazards model, respectively. We propos...
متن کاملDensity Estimation: New Spline Approaches and a Partial Review
Apart from kernel estimators, there have been quite a few different approaches of “generalized splines” for density estimation. In the present paper,Maximum Penalized Likelihood (mpl) approaches are reviewed. In conclusion, penalizing the log density seems most promising. In my “wp” approach for semi-parametric density estimation, a novel roughness penalty is introduced. It penalizes a relative...
متن کاملThe Negative Binomial Distribution Efficiency in Finite Mixture of Semi-parametric Generalized Linear Models
Introduction Selection the appropriate statistical model for the response variable is one of the most important problem in the finite mixture of generalized linear models. One of the distributions which it has a problem in a finite mixture of semi-parametric generalized statistical models, is the Poisson distribution. In this paper, to overcome over dispersion and computational burden, finite ...
متن کاملSemi-parametric Modelling of Excesses above High Multivariate Thresholds with Censored Data
One commonly encountered problem in statistical analysis of extreme events is that very few data are available for inference. This issue is all the more important in multivariate problems that the dependence structure among extremes has to be inferred. In some cases, e.g. in environmental applications, it is sometimes possible to increase the sample size by taking into account historical or inc...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید
ثبت ناماگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید
ورودعنوان ژورنال:
- Statistical applications in genetics and molecular biology
دوره 7 1 شماره
صفحات -
تاریخ انتشار 2008